Tag
28 articles
Qualcomm is designing custom chips for AWS, with a focus on AI inference, while using AWS Bedrock to optimize the chip development process.
This article explains NVIDIA's Personal AI Router (PAIR), a virtual inference router that distributes AI workloads across local devices. It covers how PAIR schedules tasks, its benefits, and its current limitations.
This article explains Perplexity's new Portable Computer system, which enables local AI inference with enhanced privacy, security, and performance through a model harness, OS-enforced sandbox, and zero per-token cost for local steps.
Liquid AI introduces LFM2.5-DSpark draft models that accelerate decoding by up to 3.18x without altering model outputs, using speculative decoding techniques.
OpenAI launches Ultrafast mode for GPT-5.6 Sol, delivering up to 750 output tokens per second using Cerebras hardware. The move introduces a three-tier pricing model centered on inference speed.
OpenAI introduces Ultrafast mode, a new API service tier that runs GPT-5.6 Sol up to 14× faster using Cerebras hardware, delivering up to 750 output tokens per second.
AMD acquires Canadian startup Taalas, which specializes in embedding AI model weights directly into silicon for ultra-fast inference. Google is reportedly pursuing a similar strategy for its Gemini models.
Google has released LiteRT.js, a JavaScript binding of its LiteRT library that enables running .tflite models in browsers via WebGPU, offering up to 60x performance gains over CPU-only execution.
OpenAI has unveiled its first custom AI processor, Jalapeño, developed in partnership with Broadcom. The chip is designed for AI inference tasks and marks a move toward vertical integration in the AI industry.
Learn how to create and run a simple AI inference example, understanding the core concepts behind AI model deployment that companies like Baseten are building upon.
Xiaomi's MiMo team, with TileRT, has achieved over 1000 tokens per second on a 1-trillion-parameter model using a single 8-GPU commodity node, marking a significant leap in LLM inference performance.
Perplexity AI introduces a hybrid local-server inference orchestrator that automatically routes AI tasks between on-device and cloud models, enhancing both performance and privacy.